Why Python? Why Colab?
Python is the language of AI. Not because it is the fastest, and not because it was designed for AI. It was not. Python is dominant in AI because of its extraordinary ecosystem of libraries: NumPy, Pandas, Scikit-learn, TensorFlow, PyTorch. Decades of brilliant people have built tools that make incredibly complex operations available in just a few lines of readable code.
Google Colab is a free, cloud-based environment that runs Python notebooks in your browser. No downloads, no installation, no "it works on my machine" problems. You get access to free GPUs, built-in libraries, and the ability to share your work with a link. It is used by researchers at DeepMind and students in their first week of learning AI alike.
print("Hello, AI world!") in a code cell and press Shift + Enter. You are now a Python programmer.The four libraries you need
You will work with four core libraries throughout this course. They are all pre-installed in Colab. This is what each one does.
import numpy as npimport pandas as pdimport matplotlib.pyplot as pltimport sklearnPython fundamentals in 5 minutes
Before we touch data, you need to know a few Python basics. These are genuinely the only things you need to get started. Do not try to memorise them. Just read them once, then use them.
# Variables — no need to declare a type name = "Alice" age = 25 is_enrolled = True score = 94.5 # Lists — ordered collections (can mix types) ages = [22, 38, 26, 35, 28] names = ["Alice", "Bob", "Charlie"] # Access elements (index starts at 0) print(ages[0]) # 22 print(ages[-1]) # 28 — last element
# For loop — iterate over a list for age in ages: print(f"Age: {age}") # If / else if age > 30: print("Over 30") elif age == 30: print("Exactly 30") else: print("Under 30") # List comprehension — a compact loop doubled = [x * 2 for x in ages] # [44, 76, 52, 70, 56]
# Define a function with def def calculate_mean(numbers): return sum(numbers) / len(numbers) mean_age = calculate_mean(ages) print(mean_age) # 29.8
Loading your first real dataset with Pandas
Now for the part that makes everything feel real. The Titanic dataset is one of the most famous datasets in machine learning. It contains information about 891 passengers, including whether each one survived. Here is how to load it and start asking questions.
import pandas as pd # Load directly from a URL — no download needed url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv" df = pd.read_csv(url) # See the first 5 rows df.head()
# Shape: how many rows and columns? print(df.shape) # (891, 12) # Statistical summary df.describe() # Average passenger age print(df['Age'].mean()) # 29.7 # Survival rate (0=died, 1=survived) print(df['Survived'].mean()) # 0.38 — 38% survived # Survival rate broken down by gender print(df.groupby('Sex')['Survived'].mean()) # female 0.742 # male 0.189
With five lines of Pandas, you have already discovered something historically significant: 74% of women survived vs 19% of men. The "women and children first" evacuation policy is visible directly in the data. This is the power of data exploration. Real signals emerge before you build a single model.
Your first plot with Matplotlib
Charts are not just pretty. They reveal patterns in seconds that would take hours to find in raw numbers. The most important habit you can build as an AI practitioner is plotting your data before doing anything else.
import matplotlib.pyplot as plt # Histogram of passenger ages plt.figure(figsize=(10, 5)) df['Age'].dropna().hist(bins=30, color='#1a6bc8', edgecolor='white') plt.title('Age Distribution of Titanic Passengers') plt.xlabel('Age') plt.ylabel('Number of Passengers') plt.show() # Survival count by class (bar chart) df.groupby('Pclass')['Survived'].mean().plot( kind='bar', color=['#2a9d8f', '#e76f51', '#264653'], title='Survival Rate by Ticket Class' ) plt.ylabel('Survival Rate') plt.xticks(rotation=0) plt.show()
Every time you call plt.show(), a chart appears directly below the cell in Colab. You will see a bell-shaped age distribution and a clear bar chart showing that 1st class passengers survived at a much higher rate than 3rd class, another stark real-world signal sitting right there in the data.
A Pandas DataFrame is like a smart spreadsheet that you can write instructions to. Instead of clicking through menus, you write short commands. df['Age'].mean() is just "give me the average of the Age column." Once you learn a few dozen of these commands, you can explore almost any dataset in the world.